Papers with reward objective

    1 papers
    Countering Reward Over-Optimization in LLM with Demonstration-Guided Reinforcement Learning (2024.findings-acl)

    Copied to clipboard

    Challenge: Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning.
    Approach: They propose a reinforcement learning approach that leverages human demonstrations and a reward model to recalibrate the reward objective.
    Outcome: The proposed approach achieves comparable performance to carefully tuned baselines while mitigating ROO in three RL language tasks.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations